Terminal-Bench 4.0 Leaderboard 2026: AI Models & Agents Ranked by Real CLI Work
Interactive Terminal-Bench 4.0 Leaderboard: GPT-6 Astra leads at 58.2% and Claude Fable 5.1 at 57.9%, with GLM-5.3 top open-weight at 41.8%. Complete scores, run costs, token counts, and verification links across 66 hard terminal tasks.
· 10 views
· Abdeladim Fadheli